In Day 16, we improved our LangGraph recommendation system by introducing conditional routing to handle the cold-start problem. However, retrieving relevant products does not necessarily mean that we have found the most suitable recommendations. Today, we will introduce LLM-based reranking by integrating Qwen through vLLM into our existing workflow. FAISS will first retrieve candidate products, and Qwen will evaluate and reorder them according to the user's preferences. Beyond improving recommendation quality, we will also introduce Codex to review our implementation. By combining a structured development harness with automated hooks, we will establish a repeatable process for inspecting code changes, running tests, and identifying potential issues before they become part of our recommendation system.
Codex writes or modifies recommendation code
↓
Harness runs validation
↓
Unit tests + linting
↓
Codex reviews changes
↓
Developer approves fixes
↓
Recommendation experiment
Harness: Your overall development process, including instructions, tests, review criteria, and validation.
Hooks: Event-triggered scripts that automate specific checks.
Codex review: AI-assisted inspection of the implementation.
1. Create an AGENTS.md file based on this templelate:
#Recommendation System Review Rules
Before completing a code change:
1. Run the existing unit tests.
2. Verify that FAISS indices map to the correct products.
3. Check that previously reviewed products are excluded.
4. Validate the cold-start routing behavior.
5. Ensure the reranker only returns candidate product IDs.
6. Check that similarity scores are not described as probabilities.
7. Report untested behavior and potential failures.
Do not modify the recommendation algorithm
without explaining the change.
2. Ask Codex to review the changes
Review my Day 17 Live recommendation system.
Focus on: 1. FAISS candidate retrieval and product-ID mapping. 2. Qwen reranking correctness and output validation. 3. Cold-start and returning-user routing. 4. Prevention of previously seen product recommendations. 5. Error handling when the vLLM server is unavailable. 6. Test coverage, latency, and maintainability.
Follow AGENTS.md. Run the available tests. Report findings by severity with file and line references. Do not modify files until I approve the fixes.

Conclusion:
Today, we extended our LangGraph recommendation system by introducing LLM-based reranking with Qwen and vLLM. Instead of relying entirely on FAISS similarity scores, our workflow first retrieves relevant candidate products and then uses Qwen to evaluate them against the user's preferences. This two-stage approach allows us to explore how semantic retrieval and LLM reasoning can work together to produce more personalized recommendations.
We also introduced Codex, harnesses, and hooks into our development workflow. By defining project-specific review rules, running automated tests, and reviewing code changes, we can identify potential issues such as incorrect product-ID mappings, invalid reranking outputs, and unexpected changes to our recommendation logic.
Reference: